Papers with dialogue metrics

3 papers
Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis (2022.naacl-main)

Copied to clipboard

Challenge: Existing frameworks for dialogue model evaluation are lacking to investigate these biases . a number of dialogue metrics are biased and can cause unforeseen problems .
Approach: They propose an adversarial test-suite which generates problematic variations of various dialogue aspects using automatic heuristics.
Outcome: The proposed test-suite generates problematic variations of various dialogue aspects using automatic heuristics.
Exploring the Impact of Human Evaluator Group on Chat-Oriented Dialogue Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Evaluator groups such as domain experts, university students, and crowdworkers have been used to assess and compare chat-oriented dialogue systems.
Approach: They analyze the impact of evaluator groups on dialogue system evaluation by testing 4 state-of-the-art dialogue systems using 4 distinct evaluer groups.
Outcome: The proposed evaluations show that the evaluator group impact is not seen for Pairwise, and that it is beneficial for certain metrics.
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities .
Approach: They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities.
Outcome: The proposed framework integrates a large number of evaluation metrics to improve the performance of the model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations